Papers with latent representations
Copied to clipboard
| Challenge: | Language modeling has advanced rapidly due to efficient model architectures and the availability of large text corpora. |
| Approach: | They propose to embed and regularize sentiment prediction-derived regularizations on the language model’s latent representations to reduce bias in the sentiment of generated text. |
| Outcome: | The proposed methods reduce bias in the sentiment of generated text by adopting individual and group fairness metrics from the fair machine learning literature. |
Copied to clipboard
| Challenge: | Using a framework of style transfer for texts, we propose several empirical methods to assess information decomposition quality. |
| Approach: | They propose to use latent representations to effectively decompose different aspects of textual information using a framework of style transfer for texts. |
| Outcome: | The proposed methods show that higher quality representations correlate with higher performance in bilingual evaluation understudy (BLEU) between output and human-written reformulations. |
Copied to clipboard
| Challenge: | Existing methods for role-playing rely on prompt engineering, which lacks stability and interpretability. |
| Approach: | They propose a framework that extracts latent representations from role-play prompts and constructs a steering vector that can be injected into the model's residual stream with controllable intensity. |
| Outcome: | The proposed framework extracts latent representations from role-play prompts, selects the most relevant features based on activation patterns, and constructs a steering vector that can be injected into the model’s residual stream with controllable intensity. |
Copied to clipboard
| Challenge: | GenerativeDictionary generates word sense interpretations based on context . traditional word sense disambiguation methods may not capture the intended word sense . |
| Approach: | They propose a dictionary system that generates word sense interpretations based on context . they transform context sentences to highlight the meaning of target words . |
| Outcome: | The proposed dictionary system is comparable to traditional word sense disambiguation methods. |
Copied to clipboard
| Challenge: | Existing approaches to index, retrieve, and read documents as evidence suffer from large computational overheads. |
| Approach: | They propose an encoder-decoder framework with an entity memory that stores entity knowledge as latent representations and pre-trained on Wikipedia along with encoder parameters. |
| Outcome: | The proposed framework outperforms memory-based and non-memory encoder-decoder models on various entity-intensive question answering and generation tasks. |
Copied to clipboard
| Challenge: | Variational Autoencoder (VAE) is an effective framework to model the interdependency for non-autoregressive neural machine translation (NAT). |
| Approach: | They propose to use Variational Autoencoder to model interdependency for non-autoregressive neural machine translation (NAT) a posterior consistency regularization approach is proposed to improve translation quality . |
| Outcome: | The proposed model is 1.5/0.7 and 0.8/0.3 BLEU points faster than the baseline model. |
Copied to clipboard
| Challenge: | Existing graph autoencoders and its variants have been used for node embedding . a new method is proposed to model consistency across different views of networks . |
| Approach: | They propose a network embedding method which enforces latent representations to be consistent across different views of networks by incorporating a multiview adversarial regularization module. |
| Outcome: | The proposed method compares favorably with the state-of-the-art methods on benchmark datasets and on a real-world application. |
Copied to clipboard
| Challenge: | Recent advances in deep learning have boosted the development of neural machine translation (NMT). |
| Approach: | They propose a flow-adapter architecture for unsupervised neural machine translation that leverages normalizing flows to model distributions of sentence-level latent representations. |
| Outcome: | The proposed model achieves competitive results on several unsupervised MT benchmarks. |
Copied to clipboard
| Challenge: | Recent models infer latent representations of words or tokens with a transformer encoder, which is bottom-up and thus does not capture long-distance context well. |
| Approach: | They propose a method to infer latent representations of words or tokens in documents . they assume a hierarchical structure of a document where top-level captures long range dependency . |
| Outcome: | The proposed model can summarize an entire book and achieve competitive performance on a wide range of document summarization benchmarks. |
Copied to clipboard
| Challenge: | Existing explanations address the contrastive aspect of explanations but their extension to textual data is under-explored and there is little investigation on their vulnerabilities and limitations. |
| Approach: | They propose a novel evaluation scheme inspired by the faithfulness of explanations by extending the computation of three metrics to textual data and benchmarking POLYJUICE and MiCE on suggested metrics. |
| Outcome: | The proposed methods demonstrate that the connectedness of counterfactuals to their original counterparts is not obvious in both models. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have impressive capabilities in natural language understanding and generation, but controlling their behavior remains a challenge. |
| Approach: | They propose a supervised steering approach that operates in sparse, interpretable representation spaces. |
| Outcome: | The proposed approach achieves higher success rates with minimal degradation in generation quality compared to existing methods. |
Copied to clipboard
| Challenge: | Euclidean space is used for training neural models and performing arithmetic operations, but many data types have complex geometries and cannot be captured in the Euclidesan space. |
| Approach: | They propose a set of guidelines for initialization, parametrization, and training of neural networks that can be generalized over existing neural network training methodologies. |
| Outcome: | The proposed framework outperforms Euclidean methods on three tasks over 12 languages and modalities on a variety of domains. |
Copied to clipboard
| Challenge: | Recent research shows that Transformer-style models can be made more efficient by sharing parameters over blocks. |
| Approach: | They propose a framework for distilling latent reasoning into a multiscale jump model that enables flexible test-time compute. |
| Outcome: | Experiments on ARC-AGI show that the proposed model achieves competitive accuracy compared to recursive baselines while requiring fewer sequential updates. |
Copied to clipboard
| Challenge: | Recent advances in large language models have demonstrated strong capabilities in tasks such as code generation and mathematical reasoning. |
| Approach: | They investigate whether large language models can construct coherent global spatial cognition by integrating fragmented relational descriptions. |
| Outcome: | The proposed models can generalize to unseen spatial relationships and exhibit latent representations aligned with real-world spatial distributions. |
Copied to clipboard
| Challenge: | Existing methods for text generation still suffer from incoherence problems . Neural sequence-to-sequence (seq2sequ) models generate fluent results . |
| Approach: | They propose a novel generation framework that leverages autoregressive self-attention mechanism to conduct content planning and surface realization dynamically. |
| Outcome: | The proposed framework outperforms baseline models and generates more coherent texts with richer contents. |
Copied to clipboard
| Challenge: | Neural models of dialog rely on generalized latent representations of language. |
| Approach: | They propose a training procedure which explicitly learns multiple representations of language at several levels of granularity. |
| Outcome: | The proposed training procedure significantly improves performance on the next utterance retrieval task using the MultiWOZ dataset and the Ubuntu dialog corpus. |
Copied to clipboard
| Challenge: | Latent multi-hop reasoning is a problem in Large Language Models that can develop shortcuts by encountering the head entity and answer entity in training sequences. |
| Approach: | They propose desiderata for shortcut-free evaluation of latent multi-hop reasoning ability . they exclude test queries where head and answer entities might have co-appeared . |
| Outcome: | The proposed model can latently recall and compose single-hop facts without shortcuts, but only for certain types of queries. |
Copied to clipboard
| Challenge: | Diffusion models have shown promise in text generation, but often struggle with generating long, coherent, and contextually accurate text. |
| Approach: | They propose a framework that enhances diffusion-based text generation through text segmentation, robust representation training with adversarial and contrastive learning, and improved latent-space guidance. |
| Outcome: | The proposed framework improves diffusion-based text generation and improves scalability and fluency. |
Copied to clipboard
| Challenge: | Existing multilingual alignment methods mitigate these issues but rely on external supervision, such as translation systems or English-biased signal. |
| Approach: | They propose a preference optimization framework that leverages an LLM’s own latent representations as intrinsic supervision signals and rewards lower-resource language outputs based on their alignment with high-resourced (English) counterparts in the "semantic hub". |
| Outcome: | The proposed framework improves a Llama 3 8B model multilingual win rates by up to 6.8% absolute (55.0% relative) on X-AlpacaEval and achieves consistent gains across benchmarks and models. |
Copied to clipboard
| Challenge: | Existing models that pretrain for cross-lingual tasks do not improve cross-linguistic learning. |
| Approach: | They propose to employ machine translation as a continued training objective to enhance language representation learning by bridging multilingual pretraining and cross-lingual applications. |
| Outcome: | The proposed model performance is compared with existing models and their latent representations. |
Copied to clipboard
| Challenge: | Conventional Deep Learning (DL)-based KT models are tied to platform-specific identifiers and latent representations, making them hard to transfer and interpret. |
| Approach: | They propose a retrieval-augmented paradigm that frames cross-platform KT as reliable context constrained inference with LLMs. |
| Outcome: | Experiments on three public KT benchmarks show that the proposed paradigm improves accuracy and robustness, and also shows strong performance under cross-platform conditions. |
Copied to clipboard
| Challenge: | Modern digital personal assistants interact with users through voice . high error rates still prevail in the widespread adoption of speech technology . |
| Approach: | They propose to extract more robust latent representations for noisy ASR text classification using transformer tokens and attentive embracement layer and multi-head attention layer. |
| Outcome: | The proposed model significantly outperforms the baseline model on the Chatbot and Snips corpora for intent classification with ASR error. |
Copied to clipboard
| Challenge: | Existing models for text generation do not need syntactic information such as constituency parses or semantic information such a paraphrase pairs. |
| Approach: | They propose a generative model which exhibits disentangled latent representations of syntax and semantics by using Attention in its decoder. |
| Outcome: | The proposed model outperforms supervised models on syntax/semantics transfer and shows that it can read latent variables with keys and values. |
Copied to clipboard
| Challenge: | Using latent optimization and Shapley values, we generate a set of minimal modifications to the text to change the classifier's prediction. |
| Approach: | They propose to generate a counterfactual by making minimal modifications to the text to change the model's prediction. |
| Outcome: | The proposed approach achieves favorable performance compared to white-box and black-box baselines using human and automatic evaluations. |
Copied to clipboard
| Challenge: | Variational autoencoders use a multivariate Gaussian latent variable to capture latent structure in data. |
| Approach: | They propose a variational autoencoder which uses a latent distribution instead of Gaussian . they find that the variational posterior averts the KL collapse by a fixed hyperparameter . |
| Outcome: | The von Mises-Fisher distribution averts the KL collapse and gives better likelihoods than Gaussian models across a range of modeling conditions. |
Copied to clipboard
| Challenge: | Syntactic structures were deemed essential in natural language processing . but since the deep learning revolution, NLP has been dominated by neural models that do not consider syntactical structures in their design. |
| Approach: | They propose a model that models latent representations of words in a sentence . they use a conditional random field to model latent and dependency arcs . |
| Outcome: | The proposed model performs competitively to transformers on small to medium sized datasets. |
Copied to clipboard
| Challenge: | Existing methods for text generation are limited in supervised setting and designed for specific applications. |
| Approach: | They propose a text generation model that learns semantics and structural features simultaneously . their model leverages a topic-based model to enhance the recognition of text semantics . |
| Outcome: | The proposed model outperforms state-of-the-art models in terms of text perplexity and topic coherence. |
Copied to clipboard
| Challenge: | Recent advances in sequence modeling have highlighted the strengths of the transformer architecture. |
| Approach: | They propose a general lattice transformer for speech translation where the input is the output of the automatic speech recognition (ASR) they propose 'controllable' lattica attention mechanism to consume latent representations. |
| Outcome: | The proposed model outperforms baseline and lattice LSTM on the Chinese-English translation task. |
Copied to clipboard
| Challenge: | a recent study shows that large language models are capable of inducing rich representations of data that are seen in-context . a novel task, adaptive world modeling, shows that even the most performant LLMs cannot reliably leverage novel semantics defined in-constitut. |
| Approach: | They propose to use in-context representations to induce rich representations of data . they also propose to probe models using a novel task to enable flexible deployment . |
| Outcome: | The proposed model can use in-context representations to complete simple downstream tasks. |
Copied to clipboard
| Challenge: | Existing latent reasoning methods that use chain of thought (CoT) are limited to selecting one discrete token at each reasoning step, which potentially induces information loss. |
| Approach: | They propose a framework that injects controllable stochasticity into latent reasoning via Gumbel-Softmax, restoring LLMs' exploratory capacity and enhancing their compatibility with Reinforcement Learning (RL). |
| Outcome: | The proposed framework preserves richer information for more comprehensive reasoning and is compatible with Reinforcement Learning (RL). |
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) is a widely adopted approach for enhancing large language models with external knowledge. |
| Approach: | They analyze how different types of retrieved documents affect the hidden states of large language models and how these internal representation shifts relate to downstream generation behavior. |
| Outcome: | The results show that context relevancy and layer-wise processing influence internal representations, providing explanations of LLMs’ output behaviors and insights for RAG system design. |
Copied to clipboard
| Challenge: | Existing knowledge-theoretic representation learning frameworks are based on the information bottleneck principle, which preserves redundant features irrelevant to the given task. |
| Approach: | They propose a conditional information flow maximization framework to learn sufficient representations for the input data and target task by maximizing both input-representation and representation-label mutual information. |
| Outcome: | The proposed framework can extract noise-invariant sufficient representations for the input data and target task. |
Copied to clipboard
| Challenge: | Existing methods for event schema generation are noise-sensitive and error-accumulating, e.g., inability to correct errors while generating schema. |
| Approach: | They propose a novel diffusion event graph model that embeds and roundes event graphs into learnable latent representations and a denoising process to maintain the model's robustness. |
| Outcome: | The proposed model achieves better results than existing state-of-the-art models on three IED bombing datasets. |
Copied to clipboard
| Challenge: | Lip reading is a process of interpreting silent speech from visual lip movements . but lip reading in cross-speaker scenarios poses a challenging problem due to inter-speech variability . |
| Approach: | They propose to exploit lip landmark-guided visual clues instead of mouth-cropped images as input features. |
| Outcome: | Experimental results show that the proposed approach reduces speaker-specific appearance characteristics in cross-speaker scenarios. |
Copied to clipboard
| Challenge: | Recent advances in mechanistic interpretability have highlighted the potential of automating interpretability pipelines in analyzing the latent representations within LLMs. |
| Approach: | They propose a framework for automatically evaluating feature-to-description alignment that measures alignment across four key metrics and quantifies the causes of misalignment. |
| Outcome: | The proposed framework evaluates alignment across four key metrics and quantifies the causes of misalignment between features and descriptions. |
Copied to clipboard
| Challenge: | Existing self-supervised learning models can learn latent representations from large amounts of unlabeled data, but they are expensive to fine-tune. |
| Approach: | They develop a meta-adapter to obtain meta-initialized parameters for self-supervised models . meta-Adapters show better generalization and extensibility than traditional pretraining methods . |
| Outcome: | Experiments on common voice and FLEURS datasets show Meta-Adapter performs better on low-resource languages . authors show it can be used on 12 low-source languages, but it requires huge computational resources . |
Copied to clipboard
| Challenge: | Existing models that map variable acoustic inputs into appropriate articulatory movements without explicit instruction are inadequate for infants. |
| Approach: | They propose a model that maps acoustic inputs into articulatory movements without explicit instruction for infants. |
| Outcome: | The proposed model outperforms MFCC features in both single- and multi-speaker settings and provides optimal representations for articulatory learning. |
Copied to clipboard
| Challenge: | Existing LLMs do not possess consistent values, but many have been developed to align them at the behavioral level, including supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF). |
| Approach: | They propose a Controlled Value Vector Activation method that directly aligns the internal values of Large Language Models by interpreting how a value is encoded in their latent representations. |
| Outcome: | The proposed method achieves highest success rate across 10 basic values without hurting model performance and fluency, and ensures target values even with opposite and potentially malicious input prompts. |
Copied to clipboard
| Challenge: | Evaluating the safety robustness of LLMs is critical for their deployment. |
| Approach: | They propose to use latent representations to characterize hidden layer dynamics by analyzing the APT of latent models and introducing the JSS metric. |
| Outcome: | The proposed method exploits the APT (Angular-Probabilistic Trajectory) of latent representations and introduces the JSS (Jensen-Shannon Separability) metric. |
Copied to clipboard
| Challenge: | Topic models aim to reveal latent structures within corpus of text through term-frequency statistics over bag-of-words representations. |
| Approach: | They propose to use bimodal vector representations of entities to extract latent representations from large language models and graph neural networks trained on symbolic relations to derive the most salient aspects of these conceptual units. |
| Outcome: | The proposed approach is better suited to working with entities than state-of-the-art models. |
Copied to clipboard
| Challenge: | A central question in multilingual language modeling is whether large language models develop a universal concept representation, disentangled from specific languages. |
| Approach: | They analyze latent representations during a word-translation task in transformer-based LLMs and extract the residual stream of the last token of the word to be translated and insert the mean at the corresponding positions in the forward pass. |
| Outcome: | The proposed model can translate a word in multiple languages without changing the language and vice versa. |
Copied to clipboard
| Challenge: | Recent advances in diffusion and conditional flow matching models for low-resolution domains are underexplored. |
| Approach: | They propose a CFM-based model that iteratively generates raw waveform in low-bitrate conditions . they propose DVQ, a factorized quantization method that uses a single quantizer . |
| Outcome: | The proposed model outperforms state-of-the-art neural audio codecs in audio quality and semantic intelligibility under low-bitrate conditions. |
Copied to clipboard
| Challenge: | Feature attribution analyses of the trained probes reveal correlations between probe accuracy and relation specificity, entity connectedness, and how distributed the signal on which the probe relies is across attention heads. |
| Approach: | They evaluate latent representations derived from attention heads and MLP contributions . they show correlations between probe accuracy and relation specificity . |
| Outcome: | The proposed representations are compared with the representations obtained from attention heads and MLPs. |